Skip to content

fix(libero): fail loudly on tasks with empty init-state sets - #267

Open
akushonkamen wants to merge 1 commit into
RLinf:mainfrom
akushonkamen:fix/246-empty-pruned-init-guard
Open

akushonkamen wants to merge 1 commit into
RLinf:mainfrom
akushonkamen:fix/246-empty-pruned-init-guard

Conversation

@akushonkamen

Copy link
Copy Markdown

Fixes #246

Description

make_env() in robots/libero/env_server.py:115-117 indexed straight into get_task_init_states(...) with no guard for tasks whose pruned_init file ships zero states:

  • seed % len(...) degenerated to seed % 0, surfacing only as a context-free ZeroDivisionError at env_server.py:117;
  • aggregation paths that tolerate the crash kept going and silently shrank the evaluation denominator (100 → 90/80), making results incomparable.

Fix:

  • Guard make_env(): when the requested task has 0 init states, raise a RuntimeError naming the suite, the requested task id, and every empty task in the suite (best-effort scan via get_num_tasks(); API signature verified against RLinf/LIBERO-PRO@rpent, liberopro/liberopro/benchmark/__init__.py:269-295).
  • Document the three currently known empty pruned_init tasks (libero_10_task#2, libero_spatial_task#3/#7) plus the HF-snapshot restore procedure in pro_hybrid_guide.md §2.3.
  • Add 3 offline fake-suite unit tests (2 for the guard, 1 healthy-task rid-math regression) in tests/unit_tests/robots/libero/test_env_server_init_states.py.

Data side (BDDL consistency check + pruned_init regeneration) belongs to the upstream LIBERO-PRO simulation pipeline per the issue triage and was not done locally; a timeout 60 probe found no LIBERO entries under ~/.cache/huggingface, so real-dataset reproduction was not executed here (the issue-sanctioned fake-suite path covers the guard).

Testing

  • Red (before fix): venvs/rlinf1662/bin/python -m pytest tests/unit_tests/robots/libero/test_env_server_init_states.py -q → 2 failed (test_make_env_fails_loudly_on_empty_init_states, test_make_env_error_lists_every_empty_task_in_suite, both ZeroDivisionError at env_server.py:117), 1 passed.
  • Green (after fix): timeout 300 venvs/rlinf1662/bin/python -m pytest tests/unit_tests -k "env_server or init_state" -q --continue-on-collection-errors → 4 passed, 1 skipped (the flag is needed because of a pre-existing, unrelated collection error in tests/unit_tests/rpent/planner/test_codex_contracts.py / openai_codex). Re-confirmed right before submission with the same numbers.
  • Regression: timeout 590 venvs/rlinf1662/bin/python -m pytest tests/unit_tests -q --continue-on-collection-errors → 14 failed, 1147 passed, 7 skipped.
  • Baseline diff (stashed fix, clean main + new test file): 16 failed = the same 14 pre-existing failures + the 2 guard tests red → zero new failures introduced by the fix.
  • Lint: timeout 120 ruff check robots/libero → All checks passed!; ruff format --check robots/libero → 18 files already formatted; tests/unit_tests/robots/libero passes both as well.
  • End-to-end: /tmp/issue246-repro/repro_zero_states.py → RuntimeError listing suite, requested task, and the full empty-task inventory for the suite.

Manual verification

Not needed — offline tests fully cover the change. The guard's runtime message was exercised end-to-end via the fake-suite repro script above (no GPU/simulator/model service involved). Real-dataset reproduction was not possible locally (no HF LIBERO cache; see Description).

Checklist

  • My code follows the code style of this project.
  • I have updated related documentation when needed.
  • I have added tests to cover my changes when needed.
  • All new and existing tests passed. (14 pre-existing failures on main are unchanged by this PR; see baseline diff above.)

AI disclosure: this change was prepared with AI assistance; the contributor has reviewed the full diff and run all listed checks.

Shipped LIBERO-PRO perturbation suites contain tasks whose pruned_init
files hold zero states (libero_10_task#2, libero_spatial_task#3/RLinf#7).
make_env() indexed straight into those sets: seed % 0 died on a cryptic
ZeroDivisionError, and any aggregation path that tolerates the crash
silently shrinks the evaluation denominator (100 -> 90/80), making
results incomparable.

Guard make_env(): when the requested task ships 0 init states, raise a
RuntimeError naming the suite, the task id and every empty task in the
suite (best-effort scan), pointing at the HF-snapshot restore path in
pro_hybrid_guide.md section 2.3. Document the currently known empty
pruned_init files there.

Offline unit tests cover the guard with a fake suite (no dataset
needed) and pin the healthy-task offset/reset-id math.

Fixes RLinf#246
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 11, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ⚠️ Failed 2026-10-11T15:19:43.411929Z 8ad9992 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Empty pruned_init files for three *_task tasks (libero_10_task t2, libero_spatial_task t3/t7)

1 participant